Papers with audio deepfake detection
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution (2025.findings-acl)
Copied to clipboard
| Challenge: | x-vector (speaker recognition PTM) achieves the highest performance in prosodic tasks . despite its low parameter, x vector captures unique prosodic characteristics of the sources . |
| Approach: | They propose to use SOTA speech pre-trained models to capture prosodic sig-natures of generative sources for audio deepfake source attribution. |
| Outcome: | The proposed model captures prosodic sig-natures of generative sources better than other models on ASVSpoof and CFAD. |
Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing algorithms for audio deepfake detection are based on layer-wise analysis of self-supervised learning (SSL) models. |
| Approach: | They conduct a layer-wise analysis of self-supervised learning (SSL) models for audio deepfake detection across diverse contexts. |
| Outcome: | The proposed models achieve competitive equal error rate (EER) scores even when employing a reduced number of layers. |
XLSR-MamBo: Scaling the Hybrid Mamba-Attention Backbone for Audio Deepfake Detection (2026.findings-acl)
Copied to clipboard
| Challenge: | Advanced speech synthesis technologies have enabled highly realistic speech generation, posing security risks that motivate research into audio deepfake detection (ADD). |
| Approach: | They propose a modular framework that integrates an XLSR front-end with synergistic Mamba-Attention backbones to capture artifacts in spoofed speech signals. |
| Outcome: | The proposed framework achieves competitive performance on the ASVspoof 2021 LA, DF, and In-the-Wild benchmarks compared to other state-of-the art systems. |